Papers with expert annotation

11 papers
Learning Interpretable Latent Dialogue Actions With Less Supervision (2022.aacl-main)

Copied to clipboard

Challenge: supervised neural dialogue modeling requires a significant amount of work to obtain turn-level labels, usually with dialogue state annotation.
Approach: They propose a novel architecture for explainable modeling of task-oriented dialogues with discrete latent variables to represent dialogue actions.
Outcome: The proposed model outperforms previous approaches with less supervision in terms of perplexity and BLEU on three datasets.
Towards Self-Improving Error Diagnosis in Multi-Agent Systems (2026.findings-acl)

Copied to clipboard

Challenge: Existing diagnostic approaches rely on expensive expert annotations and ”LLM-as-a-judge” paradigms.
Approach: They propose a framework for semantic failure attribution that identifies responsible agents and the originating error step.
Outcome: The proposed framework outperforms baselines in step-level localization and validation.
A Probabilistic Annotation Model for Crowdsourcing Coreference (D18-1)

Copied to clipboard

Challenge: Existing methods to generate annotated corpora for coreference are expensive and limited.
Approach: They propose a model of annotation for aggregating crowdsourced anaphoric annotations.
Outcome: The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation.
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in multimodal large language models (MLLMs) offer new opportunities for higher-level scene understanding, but they require labor-intensive, expert annotation.
Approach: They propose a dataset that combines 2K human-verified images with 22K image-description pairs to provide a more accurate representation of pedestrian scenes.
Outcome: The proposed dataset improves scalability while maintaining quality.
How coherent are neural models of coherence? (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to model coherence are limited to small newswire corpora . evaluators need to be trained on lexical and document levels to perform evaluations .
Approach: They propose four generic evaluation tasks that capture coherence-specific properties . they aim at capturing correct use of discourse connectives and lexical cohesion .
Outcome: The proposed tasks capture coherence-specific properties, including correct use of discourse connectives, lexical cohesion, temporal consistency among events and participants in a story.
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Recent research emphasizes the generation of high-quality feedback that provides justification and actionable guidance.
Approach: They propose an LLM-based framework for evaluating LLM feedback along three dimensions: specificity, helpfulness, and validity.
Outcome: The proposed framework evaluates LLM-generated feedback along three dimensions: specificity, helpfulness, and validity.
DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging (2024.emnlp-main)

Copied to clipboard

Challenge: Modern large language models (LLMs) showcase impressive capabilities across various tasks with aligning their behavior with human preferences.
Approach: They propose a framework that integrates domain-specific knowledge into a general reward model by model merging.
Outcome: The proposed framework improves performance across different benchmarks and provides detailed analysis showing the effects of model merging.
LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval (2025.emnlp-main)

Copied to clipboard

Challenge: Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is limited by incomplete or inconsistent textual descriptions.
Approach: They propose a data augmentation framework that enhances cross-modal retrieval performance by improving the completeness and consistency of LLM-generated descriptions.
Outcome: The proposed framework improves cross-modal retrieval performance by improving completeness and consistency of LLM-generated descriptions.
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.
Approach: They propose to use an ensemble of large language models to flag mislabeled examples by using an LLM-as-a-judge approach to detect label errors in existing datasets.
Outcome: The proposed method improves label accuracy and consistency in large language models.
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Current systems for legal consultation are insufficient to handle the knowledge-intensive nature of real-world consultations.
Approach: They propose a multi-turn benchmark dataset to evaluate LLMs in legal consultation settings.
Outcome: The proposed framework assesses LLMs’ consultation capabilities in terms of (1) clarification capability and (2) professional advice quality.
XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics (2026.findings-acl)

Copied to clipboard

Challenge: averaging metric scores across languages is suspicious since translations of equal quality receive different scores across language.
Approach: They propose a semi-automatically built dataset to benchmark translation metrics using MQM-defined errors and a normalization strategy to mitigate cross-lingual scoring bias.
Outcome: The proposed model shows that translation metrics suffer from cross-lingual scoring bias . the proposed model is based on a semi-automatically built dataset covering nine translation directions .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations